On Explore-Then-Commit strategies
نویسندگان
چکیده
We study the problem of minimising regret in two-armed bandit problems with Gaussian rewards. Our objective is to use this simple setting to illustrate that strategies based on an exploration phase (up to a stopping time) followed by exploitation are necessarily suboptimal. The results hold regardless of whether or not the difference in means between the two arms is known. Besides the main message, we also refine existing deviation inequalities, which allow us to design fully sequential strategies with finite-time regret guarantees that are (a) asymptotically optimal as the horizon grows and (b) order-optimal in the minimax sense. Furthermore we provide empirical evidence that the theory also holds in practice and discuss extensions to non-gaussian and multiple-armed case.
منابع مشابه
Computing Optimal Strategies to Commit to in Stochastic Games
Significant progress has been made recently in the following two lines of research in the intersection of AI and game theory: (1) the computation of optimal strategies to commit to (Stackelberg strategies), and (2) the computation of correlated equilibria of stochastic games. In this paper, we unite these two lines of research by studying the computation of Stackelberg strategies in stochastic ...
متن کاملThe Role of Critical Thinking Orientation on the Learners’ Use of Communicative Strategies
Abstract The present study aimed to explore the impact of teaching critical thinking skills through applying debate on the use of communicative strategies. At first 60 intermediate students were selected and placed in two homogenous groups of control and experimental through passing Nelson test. Then, a critical thinking appraisal was run to the two groups both before and after the treatment. T...
متن کاملThe Role of Critical Thinking Orientation on the Learners’ Use of Communicative Strategies
Abstract The present study aimed to explore the impact of teaching critical thinking skills through applying debate on the use of communicative strategies. At first 60 intermediate students were selected and placed in two homogenous groups of control and experimental through passing Nelson test. Then, a critical thinking appraisal was run to the two groups both before and after the treatment. T...
متن کاملDisarmament Games
Much recent work in the AI community concerns algorithms for computing optimal mixed strategies to commit to, as well as the deployment of such algorithms in real security applications. Another possibility is to commit not to play certain actions. If only one player makes such a commitment, then this is generally less powerful than completely committing to a single mixed strategy. However, if p...
متن کاملOptimal Machine Strategies to Commit to in Two-Person Repeated Games
The problem of computing optimal strategy to commit to in various games has attracted intense research interests and has important real-world applications such as security (attacker-defender) games. In this paper, we consider the problem of computing optimal leader’s machine to commit to in two-person repeated game, where the follower also plays a machine strategy. Machine strategy is the gener...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2016